Papers with fusion strategies

11 papers
Ambiguity-aware Multi-level Incongruity Fusion Network for Multi-Modal Sarcasm Detection (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for sarcasm detection focus on fusing text and image information to establish cross-modal correlations, overlooking the significance of original unimodal incongruity information.
Approach: They propose a multi-modal incongruity learning module to capture inconcluity information simultaneously at the text-level, image-level and cross-modal-level.
Outcome: The proposed model outperforms state-of-the-art methods on a publicly available dataset.
Using Multi-Encoder Fusion Strategies to Improve Personalized Response Selection (2022.coling-1)

Copied to clipboard

Challenge: Existing systems that focus on persona do not explore well the correlation between persona and empathy.
Approach: They propose a suite of fusion strategies that capture interaction between persona, emotion, and entailment information of the utterances.
Outcome: The proposed model outperforms the previous methods by 2.3% on original personas and 1.9% on revised persona models in terms of hits@1 accuracy.
Multimodal Quality Estimation for Machine Translation (2020.acl-main)

Copied to clipboard

Challenge: Existing work has only explored textual context.
Approach: They propose to use visual and text modalities to explore Quality Estimation for Machine Translation and integrate them into multimodal QE frameworks.
Outcome: The proposed approaches improve on sentence-level and document-level predictions using visual features extracted from images.
Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level Alignment (P18-1)

Copied to clipboard

Challenge: Existing approaches to classify human affect and subjective information from multiple data sources are limited by the lack of high-level feature associations.
Approach: They propose a hierarchical multimodal architecture with attention and word-level fusion to classify utterance-level sentiment and emotion from text and audio data.
Outcome: The proposed model outperforms state-of-the-art approaches on published datasets and visualizes and interprets synchronized attention over modalities.
Uncertainty-Guided Modal Rebalance for Hateful Memes Detection (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for integrating hate information from different modalities ignore the modality uncertainty caused by the contribution degree of each modality to hate sentiment.
Approach: They propose an Uncertainty-guided Modal Rebalance framework for hateful memes detection . they propose to combine cross-modal fusion features with unimodal features .
Outcome: The proposed framework produces state-of-the-art performance on four widely-used datasets.
MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown impressive progress in mathematical problem-solving . current approaches to enhance mathematical reasoning focus on instance-level modifications .
Approach: They propose a framework that enhances mathematical reasoning through cross-problem instruction synthesis.
Outcome: The proposed framework boosts mathematical reasoning by 18.0 points while maintaining high data efficiency.
Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 Videos (2022.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal summarization methods are limited to monolingual videos . a proposed task aims to generate cross-lingual summaries from multimodal inputs .
Approach: They propose a task to generate cross-lingual summaries from multimodal inputs of videos . they propose fusion network that integrates multimodal and cross-linguistic information .
Outcome: The proposed task outperforms existing methods on a reorganized How2 dataset on the reorganized How2 data set.
What to Fuse and How to Fuse: Exploring Emotion and Personality Fusion Strategies for Explainable Mental Disorder Detection (2023.findings-acl)

Copied to clipboard

Challenge: Mental health disorders (MHD) are one of the greatest challenges facing our healthcare systems and modern societies in general.
Approach: They integrate and extend the research by conducting extensive experiments with three types of deep learning-based fusion strategies: feature-level fusion, model fusion and task fusion.
Outcome: The proposed techniques show that they can be used to improve mental health detection from textual data.
Latent Distribution Decouple for Uncertain-Aware Multimodal Multi-label Emotion Recognition (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on improving fusion strategies and modeling modality-to-label dependencies, but they overlook the impact of aleatoric uncertainty, which is inherent noise in multimodal data.
Approach: They propose a latent emotional distribution decomposition with uncertainty perception framework to model aleatoric uncertainty in multimodal data.
Outcome: The proposed framework achieves state-of-the-art performance on the CMU-MOSEI and M3ED datasets, highlighting the importance of uncertainty modeling in MMER.
A Multi-View Media Profiling Suite: Resources, Evaluation, and Analysis (2026.findings-acl)

Copied to clipboard

Challenge: a large-scale label set for media outlets from Media Bias/Fact Check (MBFC) is lacking in the field.
Approach: They propose to use a large-scale label set to analyze outlets' representations . they also propose to evaluate embedding views and fusion strategies .
Outcome: The proposed method achieves state-of-the-art results on ACL-2020 and establishes strong benchmarks on MBFC-2025.
Recurrent Knowledge Identification and Fusion for Language Model Continual Learning (2025.acl-long)

Copied to clipboard

Challenge: Continual learning (CL) is crucial for large language models without costly retraining.
Approach: They propose a framework for recurrent knowledge identification and fusion that enables dynamic estimation of parameter importance distributions to enhance knowledge transfer.
Outcome: The proposed framework mitigates catastrophic forgetting and enhances knowledge transfer.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations